Papers with comprehension of Reichenbach
TimeRes: A Turkish Benchmark For Evaluating Temporal Understanding of Large Language Models (2026.eacl-srw)
Copied to clipboard
| Challenge: | Existing benchmarks focus on English and underexplore how linguistic structure contributes to temporal meaning. |
| Approach: | They propose a Turkish benchmark to evaluate temporal understanding of Large Language Models (LLMs) their benchmark examines Reichenbach’s temporal points and reported speech through date arithmetic . |
| Outcome: | The proposed model fails to resolve reported speech and fails to generalize across word order variations. |